Papers with image captioning models

15 papers
Combine to Describe: Evaluating Compositional Generalization in Image Captioning (2022.acl-srw)

Copied to clipboard

Challenge: Recent work on compositionality has focused on the ability to combine simpler concepts to understand & generate arbitrarily more complex conceptual structures.
Approach: They propose to use a set of image captioning models to benchmark their compositional generalization properties.
Outcome: The proposed models do not generalize in terms of systematicity and productivity, but are robust to synonym substitutions.
Fine-grained Image Captioning with CLIP Reward (2022.findings-naacl)

Copied to clipboard

Challenge: Modern image captioning models are usually trained with text similarity objectives . reference captions often describe only the most salient objects in images .
Approach: They propose to use CLIP to calculate multi-modal similarity and use it as a reward function . they propose a simple finetuning strategy to improve grammar that does not require extra text annotation.
Outcome: The proposed model generates more distinctive captions than the CIDEroptimized model on text-to-image retrieval and fineCapEval.
Are Scene Graphs Good Enough to Improve Image Captioning? (2020.aacl-main)

Copied to clipboard

Challenge: Existing image captioning models rely on object detection features to generate image descriptions, but they are noisy.
Approach: They propose to use scene graphs to introduce information about object relations into captioning to improve image descriptions.
Outcome: The proposed model improves image caption quality by 3.3 CIDEr compared to a strong Bottom-Up Top-Down baseline.
The Role of Data Curation in Image Captioning (2024.eacl-long)

Copied to clipboard

Challenge: Existing image captioning models treat all samples equally, neglecting mismatched data . Several other techniques have relied on curriculum learning strategies to adapt learning to the difficulty of the task.
Approach: They propose to actively curate difficult samples in datasets using curriculum learning strategies to improve captioning models.
Outcome: The proposed methods outperform existing models on the Flickr30K and COCO datasets.
What Makes for Good Image Captions? (2025.findings-emnlp)

Copied to clipboard

Challenge: a formal information-theoretic framework is developed for image captioning . the pyramid of captions is a method that generates enriched captions by integrating local and global visual information.
Approach: They propose a formal information-theoretic framework for image captioning . they propose 'Pyramid of Captions' method that generates enriched captions .
Outcome: The proposed framework provides a flexible foundation for analyzing and optimizing image captioning systems across diverse task requirements.
MICE: Mixture of Image Captioning Experts Augmented e-Commerce Product Attribute Value Extraction (2025.acl-industry)

Copied to clipboard

Challenge: Existing visual attribute value extraction methods rely on product images and textual information, which can be ambiguous, inaccurate, or unavailable.
Approach: They propose a framework that leverages a curated pool of image captioning models to generate accurate captions from product images.
Outcome: The proposed framework significantly improves state-of-the-art large multimodal models in zero-shot and fine-tuning settings.
Training for Diversity in Image Paragraph Captioning (D18-1)

Copied to clipboard

Challenge: Existing image captioning models have a lack of diversity between sentences . current models have limited their effectiveness due to repetitive paragraphs .
Approach: They propose to apply sequence-level training to image paragraph captioning models . they find that standard self-critical training produces poor results .
Outcome: The proposed training improves on the Visual Genome dataset with no architectural changes.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning (P18-1)

Copied to clipboard

Challenge: Practical applications of automatic image description systems include leveraging descriptions for image indexing or retrieval, and helping those with visual impairments by transforming visual signals into information that can be communicated via text-to-speech technology.
Approach: They propose to extract and filter image caption annotations from billions of webpages and use them to train models.
Outcome: The proposed model architectures perform better when trained on the Conceptual Captions dataset.
Transparent Human Evaluation for Image Captioning (2022.naacl-main)

Copied to clipboard

Challenge: Recent work has demonstrated that image captioning is a complex task that requires a large amount of human input.
Approach: They develop a human evaluation protocol for image captioning models based on machine- and human-generated captions on the MSCOCO dataset.
Outcome: The proposed model improves CLIPScore, a recent metric that uses image features, and improves human judgments because it is more sensitive to recall.
Bridge the Gap: High-level Semantic Planning for Image Captioning (2020.coling-main)

Copied to clipboard

Challenge: Recent image captioning models have improved the multi-modal interaction, such as attention mechanisms.
Approach: They propose a high-level semantic planning mechanism that integrates a semantic reconstruction and an explicit order planning mechanism to bridge the gap between visual and language domains.
Outcome: The proposed model outperforms previous methods and achieves the state-of-the-art performance on MS COCO.
Object Hallucination in Image Captioning (D18-1)

Copied to clipboard

Challenge: Existing image captioning metrics do not capture image relevance . current metrics only measure similarity to ground truth captions .
Approach: They propose a new image relevance metric to evaluate captioning models with veridical visual labels and assess their rate of object hallucination.
Outcome: The proposed metrics show that models with veridical visual labels have higher hallucination rates than models with lower hallucinosity.
IC3: Image Captioning by Committee Consensus (2023.emnlp-main)

Copied to clipboard

Challenge: Traditionally, image captioning models are trained to generate a single “best’ (most like a reference) image caption.
Approach: They propose a method to generate a single caption that captures high-level details from several annotator viewpoints.
Outcome: The proposed method outperforms baseline SOTA models and improves the performance of automated recall systems by up to 84%.
PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary Representations (2024.acl-long)

Copied to clipboard

Challenge: Existing models for word sense disambiguation lack images or senses in textual and visual datasets.
Approach: They propose a unified image-text WSD model that uses image-sense complementarity to generate visual representations for word senses and a disambiguation-oriented image-sensor dataset to provide implicit textual representations.
Outcome: The proposed model achieves 2.53% F1-score increase over state-of-the-art models on Textual-WSD and 2.22% HR@1 improvement on Visual-WSS.
Cross-modal Coherence Modeling for Caption Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for image captioning do not guarantee consistent image-text relations . current models do not provide enough data for training robust captioning models .
Approach: They use an annotation protocol specifically devised for capturing image–caption coherence relations to study image captioning.
Outcome: The proposed protocol improves image captioning models with coherence relations . the dataset is large enough to alleviate content hallucinations, the authors show .
Mitigating Open-Vocabulary Caption Hallucinations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for image captioning ignore the long-tailed nature of hallucinations . a new framework is proposed to address hallucines in image captions in the open-vocabulary setting .
Approach: They propose a framework to address hallucinations in image captioning in the open-vocabulary setting.
Outcome: The proposed framework surpasses the CHAIR benchmark in diversity and accuracy in open-vocabulary captioning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations